Skip to content

docs: timeout-minutes was not what kills the sweep — I was wrong - #86

Merged
LMPrado-DZ23 merged 1 commit into
release/v3.8.55from
docs/sweep-60min-finding
Sep 20, 2026
Merged

LMPrado-DZ23 merged 1 commit into
release/v3.8.55from
docs/sweep-60min-finding

Conversation

@LMPrado-DZ23

Copy link
Copy Markdown
Owner

#76 gave both sweep jobs timeout-minutes: 180 on the theory that the job, declaring no budget, inherited one. That theory was wrong, and the next run disproved it.

Run 35496465106 was dispatched on 450d6586f — which already carried the 180 — and died at the same mark as the three attempts before it:

07:19:01  ✅ [hard]  Typecheck (core)
07:22:03  ✅ [hard]  ESLint errors — 0
07:22:03  ✅ [drift] ESLint warnings — 0 vs baseline 1050
07:22:39  ✅ [drift] Cognitive complexity — 1234 vs baseline 1437
07:23:12  ✅ [drift] Type coverage, Compression budget, OpenAPI coverage,
                     Workflow lint, CodeQL — all within baseline
07:23:23  ✅ [hard]  Docs sync + fabricated-docs (strict)
          ▶ the four serial suites start
08:17:37  ##[error] Process completed with exit code 143

Job start 07:17:26 → kill 08:17:37. 60 minutes and 11 seconds.

What it actually is

Every static and drift gate passes. The suites then ran for 54 minutes with no output and the process took SIGTERM. Run 35468579833 had recorded the reason in words all along, and I did not weigh it properly:

The runner has received a shutdown signal

The hosted runner stops at ~60 minutes. The suites' own ceilings — 80 (unit) + 15 (vitest) + 40 (integration) + 20 (pack) = 155 minutes worst case in serial — cannot fit inside that.

Two real options, neither reachable by editing a workflow

  • set the USE_VPS_RUNNER repository variable for a release window: the workflow already honours it and it moves the sweep off the hosted runner;
  • shard the slow suites across jobs so no single job needs more than an hour.

What this PR changes

Nothing executable. The timeout-minutes lines stay — an explicit budget is still better than an inherited one, and a sweep that overruns that reports a timeout instead of an opaque 143. What changes is the comment: it claimed to be the cause, and now it records that it is not, with the measurement, so the next person does not repeat my attempt. audit/FINAL_THREE_AGENT_REVIEW.md carries the same correction.

The architecture audit's HIGH stays open. No full-CI verdict exists for this line. What changed is that the reason is now known and written down instead of being guessed at a fourth time.

Gate Result
YAML parses, both sweep jobs still at 180 ✓
check:doc-links PASS
prettier --check clean

🤖 Generated with Claude Code

#76 gave both sweep jobs `timeout-minutes: 180` on the theory that the
job, declaring no budget, inherited one. Run 35496465106 disproved it.

That run was dispatched on 450d658, which ALREADY carried the 180, and
it died at 08:17:37 — 60m11s after the job started, the same mark as the
three attempts before it:

  07:19:01  ✅ [hard]  Typecheck (core)
  07:22:03  ✅ [hard]  ESLint errors — 0
  07:22:39  ✅ [drift] Cognitive complexity — 1234 vs baseline 1437
  07:23:23  ✅ [hard]  Docs sync + fabricated-docs (strict)
            ▶ the four serial suites start
  08:17:37  ##[error] Process completed with exit code 143

Every static and drift gate passes. The suites then ran 54 minutes with
no output and the process took SIGTERM. Run 35468579833 had recorded the
reason in words all along: "The runner has received a shutdown signal".

So the hosted runner stops at ~60 minutes, and the suites' own ceilings
(80 + 15 + 40 + 20 = 155 min worst case, serial) cannot fit inside it.
Two real options, neither reachable by editing a workflow:

  - set USE_VPS_RUNNER for a release window — the workflow already
    honours it and it moves the sweep off the hosted runner;
  - shard the slow suites so no single job needs more than an hour.

The `timeout-minutes` lines stay: an explicit budget is still better
than an inherited one, and a sweep that overruns THAT reports a timeout
instead of an opaque 143. But the comment claiming it was the cause is
now a comment saying it is not, with the measurement, so nobody repeats
my attempt.

audit/FINAL_THREE_AGENT_REVIEW.md carries the same correction — the HIGH
stays open, and what changed is that the reason is known.

  YAML parses, both sweep jobs still at 180; prettier clean
  [doc-links] PASS

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 20, 2026

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 3ef448c3-81ca-482e-a1a1-13909253b96c


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@LMPrado-DZ23
LMPrado-DZ23 merged commit 9cebfd4 into release/v3.8.55 Sep 20, 2026
15 checks passed
@LMPrado-DZ23
LMPrado-DZ23 deleted the docs/sweep-60min-finding branch September 20, 2026 08:47
LMPrado-DZ23 pushed a commit that referenced this pull request Sep 20, 2026
Four consecutive full sweeps died at exit 143 and none of them produced a
verdict for release/v3.8.55. #76 blamed a missing `timeout-minutes`; #86
records the measurement that disproved it. The real shape:

  07:17:26  job starts
  07:19:01  ✅ Typecheck   07:22:03 ✅ ESLint   07:22:39 ✅ complexity
  07:23:23  ✅ Docs sync + fabricated-docs — every static gate green
            ▶ four serial suites start
  08:17:37  ##[error] exit 143        ← 60m11s

A runner on this repository stops at ~60 minutes ("The runner has received a
shutdown signal", run 35468579833). The suites' own ceilings are 80 + 15 + 40
+ 20 = 155 minutes serial. No single job can carry that, and no value of
`timeout-minutes` changes it. So the sweep stops being one job.

  resolve      → resolves the branch ONCE and outputs the exact SHA
  slow-suite   → 7 parallel jobs: unit ×4, integration ×2, vitest
  release-green → static + drift + full-ci + pack + boot, then MERGES their
                  reports and owns the verdict

Every job checks out the SHA `resolve` produced, so the merged report belongs
to one commit rather than to whatever each job happened to fetch.

The shards measure; they do not judge. A red suite still exits 0 so its report
reaches the aggregator — and the aggregator runs on `!cancelled()`, not
`success()`, so a shard that DIED still gets its verdict pronounced.

That last part is the whole risk of this shape, so it is the part with teeth:

  --expect-slow names every shard id the matrix produces. A suite with no
  report is recorded as a HARD failure reading "it did not run, so it is NOT
  green". Not a warning, not an absence.
  --slow-gates=<typo> throws instead of selecting nothing (a job that runs
  zero tests must not report green).
  A run that records zero gates is a HARD failure for that reason alone.
  A test derives the expected ids FROM the matrix and compares; it fails
  closed, so a stale list goes red rather than quiet.

This is the same defect class I have now fixed three times in this branch —
a gate reporting green while measuring nothing — so it is guarded before it
can happen rather than after.

  47/47 tests/unit/validate-release-green.test.ts
  red-first: dropping two unit shards from the expect list fails test 46
  --no-static --slow-gates=none → ❌ "ran zero gates"
  --slow-gates=unit-tests       → refuses to start
  --merge-slow without --expect-slow → refuses to start
  merge smoke: 2 unit shards green + integration absent → ❌ NOT release-green
  YAML parses; prettier clean

Still open, and not touched here: `main-green` carries the same 155 minutes in
one job and will keep dying at 60 minutes. It validates the released state,
not the release candidate, so it does not block this line — but it is the same
fix, and it is not done.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
LMPrado-DZ23 added a commit that referenced this pull request Sep 20, 2026
* ci: split the release-green sweep so a hosted runner can finish it

Four consecutive full sweeps died at exit 143 and none of them produced a
verdict for release/v3.8.55. #76 blamed a missing `timeout-minutes`; #86
records the measurement that disproved it. The real shape:

  07:17:26  job starts
  07:19:01  ✅ Typecheck   07:22:03 ✅ ESLint   07:22:39 ✅ complexity
  07:23:23  ✅ Docs sync + fabricated-docs — every static gate green
            ▶ four serial suites start
  08:17:37  ##[error] exit 143        ← 60m11s

A runner on this repository stops at ~60 minutes ("The runner has received a
shutdown signal", run 35468579833). The suites' own ceilings are 80 + 15 + 40
+ 20 = 155 minutes serial. No single job can carry that, and no value of
`timeout-minutes` changes it. So the sweep stops being one job.

  resolve      → resolves the branch ONCE and outputs the exact SHA
  slow-suite   → 7 parallel jobs: unit ×4, integration ×2, vitest
  release-green → static + drift + full-ci + pack + boot, then MERGES their
                  reports and owns the verdict

Every job checks out the SHA `resolve` produced, so the merged report belongs
to one commit rather than to whatever each job happened to fetch.

The shards measure; they do not judge. A red suite still exits 0 so its report
reaches the aggregator — and the aggregator runs on `!cancelled()`, not
`success()`, so a shard that DIED still gets its verdict pronounced.

That last part is the whole risk of this shape, so it is the part with teeth:

  --expect-slow names every shard id the matrix produces. A suite with no
  report is recorded as a HARD failure reading "it did not run, so it is NOT
  green". Not a warning, not an absence.
  --slow-gates=<typo> throws instead of selecting nothing (a job that runs
  zero tests must not report green).
  A run that records zero gates is a HARD failure for that reason alone.
  A test derives the expected ids FROM the matrix and compares; it fails
  closed, so a stale list goes red rather than quiet.

This is the same defect class I have now fixed three times in this branch —
a gate reporting green while measuring nothing — so it is guarded before it
can happen rather than after.

  47/47 tests/unit/validate-release-green.test.ts
  red-first: dropping two unit shards from the expect list fails test 46
  --no-static --slow-gates=none → ❌ "ran zero gates"
  --slow-gates=unit-tests       → refuses to start
  --merge-slow without --expect-slow → refuses to start
  merge smoke: 2 unit shards green + integration absent → ❌ NOT release-green
  YAML parses; prettier clean

Still open, and not touched here: `main-green` carries the same 155 minutes in
one job and will keep dying at 60 minutes. It validates the released state,
not the release candidate, so it does not block this line — but it is the same
fix, and it is not done.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): fragment for #87

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci(main-green): say plainly that this job still cannot finish

The comment on main-green was left describing sharding as a future
option, which is no longer true of the release path above — it IS
sharded now — and it still carried `timeout-minutes: 180`, the number
run 35496465106 disproved.

main-green is knowingly NOT fixed. It is off by default
(VALIDATE_MAIN_BRANCH) and this fork ships from release/v* and never
merges into main, so main is a stale upstream snapshot nobody here
releases. Splitting it would duplicate ~80 lines of matrix for a job
that does not run.

What changes is the failure mode. 55 minutes rather than 180 means
whoever enables it gets a stated timeout with the log, instead of the
opaque exit 143 that four consecutive full sweeps died of — and the
comment now names the fix (the slow-suite matrix above) instead of
describing it as unreachable.

The budget test covers all three jobs now, main-green included.

  47/47 tests/unit/validate-release-green.test.ts
  YAML parses, main-green timeout: 55; prettier clean

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(ops): document the shard flags in RELEASE_GREEN.md

The sweep's new shape is only usable if the flags that split it are
written down next to the ones they extend — and the three refusals that
keep a split verdict honest (--expect-slow, unknown suite, zero gates)
belong with them, not only in the script's header.

  [doc-links] PASS — 172 docs, 1044 internal links
  fabricated-claim gate: clean

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: zodyp <zodyprado@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants